The goal of this study was to investigate the potential impact of incorporating methods for detecting IER into questionnaires commonly used in applied psychological research. We conducted an experiment comparing three methods recommended in the literature (bogus items, instructed response items, and instructed manipulation checks), either preceded or not preceded by a warning, to a control questionnaire. Participants in the experiment were working adults in Austria, Germany, and Switzerland. As our sample is thus relatively close to target populations typically studied in applied psychological research, particularly in comparison to some earlier investigations of IER, we submit that our findings should generalise relatively well to “real-world” applications of IER detection methods. For data analysis, we adopted an invariance testing approach to assess whether IER detection influenced response behaviours in a way that would affect estimates of substantive interest to researchers. We submit that the findings from our study provide three important insights, which we discuss below.
First, our findings suggest that overall, the use of IER detection methods has relatively little impact on participants' response behaviour and most resulting parameter estimates. Specifically, we found that IER detection methods have no significant impact on factor loadings and item intercepts—two important parts of the relationship between measurement items and their common factor. Evidently, study participants tended to interpret substantive items in a questionnaire in a very similar way, regardless of whether it contained IER detection methods (but see our discussion of reliability below). Similarly, with respect to parameters that are typically of greatest interest to substantive researchers, we found that IER detection methods have little impact on factor means and factor covariances. Thus, estimates of group means (as in, e.g., analysis of variance) or relationships between variables (as in, e.g., correlation, regression, or path modelling) appear not to be affected by the presence of IER detection methods. Importantly, the reaction measures (Huang, Bowling et al., 2015) we employed to gauge participants' perceptions of the questionnaire were also unaffected. This finding supports earlier research suggesting that using IER detection methods does not deteriorate the experience of study participation (Huang, Bowling et al., 2015), and may be interpreted as some evidence against the notion of mood being a driver of differential survey responses (Tourangeau et al., 2009). We do note, however, that there was some evidence in our study suggesting that using IER detection methods might affect estimates of factor variances. We hesitate to consider this evidence strong enough to permit clear recommendations. Further, investigating the invariance of factor variance estimates in the context of IER detection presents an opportunity for future research.
Second, our study did provide evidence for estimates of reliability (i.e. internal consistency) being affected by IER detection methods. Specifically, we found that questionnaires using such methods without alerting respondents to their presence produced decreased reliability estimates (Schmitt & Kuljanin, 2008), albeit not for all constructs. Yet, given that reliability estimates in our study accounted for measurement error (Cho, 2016), some of the more severely reduced estimates (in comparison to the control group; see Table 5) may be considered problematic with respect to the standards proposed for reliability in applied psychological research (Greco, O'Boyle, Cockburn, & Yuan, 2016). This is because unreliability can threaten both the statistical conclusion validity (because estimated relationships between variables may be biased in either direction) and the construct validity (due to biased estimates of relationships with other variables in the nomological network) of a study (Shadish, Cook, & Campbell, 2002). As researchers are routinely expected to assess and report reliability of their measures, we argue that these issues should be carefully considered, even if the variables in question are “merely” intended for use as control variables (Aguinis & Vandenberg, 2014). Investigating why reliability was reduced by some IER detection methods (i.e. the psychological processes in respondents; e.g. Schwarz, 1999; Tourangeau et al., 2009) was beyond the scope of this study, and presents opportunities for future research. This also includes elucidating the boundary conditions under which IER detection methods may act as deterrents (Hauser & Schwarz, 2015; Miller & Baker-Prewitt, 2009), potentially increasing reliability as more attentive respondents tend to exhibit measurements with higher reliability (Maniaci & Rogge, 2014).
Interestingly, our study also found that including a warning informing respondents of the use of IER detection methods produced relatively higher reliabilities of measures. This would be consistent with the notion that warnings may serve as a tool for deterring IER (Huang et al., 2012; Huang, Bowling et al., 2015) by increasing attentiveness which, again, may increase reliability (Maniaci & Rogge, 2014). This, in combination with our finding that respondents in those same conditions tended to report slightly more positive reactions to the questionnaire (Huang, Bowling et al., 2015), suggests that the potentially adverse impact of using IER detection methods may be counteracted by adding a simple warning to the questionnaire. Therefore, we recommend that researchers incorporating IER detection methods into their questionnaires should consider also alerting respondents to this very fact.
Third and finally, our study provides some insights into the relative “performance” of IER detection methods. We found that, compared to IR items and IMCs, bogus items tended to have relatively smaller impact on parameter estimates across types of parameters (including error variances and thus reliability; see above). In addition, the use of bogus items had virtually no impact, particularly no adverse impact, on respondent reactions to participating in the survey (although IR items and IMCs performed similarly). Again, these findings bolster earlier research suggesting that bogus items may be a useful tool for detecting IER (Huang, Bowling et al., 2015). Moreover, while bogus items did not differ dramatically from other methods in any single aspect, overall results suggest that they might have a slight advantage over IR items and IMCs.
We note several limitations of our study. Our findings on bogus items are limited in their generalisability to other bogus items. This is because, among the bogus items tested in the extant literature (Dunn et al., in press; Huang, Bowling et al., 2015; Meade & Craig, 2012), those included in our questionnaire were arguably relatively “mild” in terms of their potential to upset, offend, or confuse study participants—which is precisely why we chose them. For example, consider the item “I have never used a computer” (which we used) versus the item “I can teleport across time and space” (Huang, Bowling et al., 2015). We would argue that, while the former might appear unusual to respondents, the latter may be perceived as more outlandish and therefore potentially more evocative of an affective reaction that may impact subsequent responses. Thus, while our findings provide some support for the use of bogus items, our study cannot speak to the usefulness of all bogus items proposed in the literature. Investigating the relative impact of different kinds of bogus items, particularly with respect to affective (Tourangeau et al., 2009; but also cognitive) reactions, presents another opportunity for future research. Our findings are also limited in terms of their comparability to other studies of IER to the extent that those studies used different samples and/or questionnaire designs and substantive measures. With respect to sampling, our sample consisted of working adults as opposed to students or MTurk workers, where the generalisability of findings from the latter is not always entirely clear (e.g. Cheung, Burns, Sinclair, & Sliter, 2016; Peterson, 2001). However, as participants in our study were members of an online respondent panel (albeit without being compensated), these individuals may still have some experience in participating in research studies. While members in this panel are restricted to four study participations per year, we cannot rule out that this may have impacted their response behaviour. Finally, with respect to questionnaire design, some earlier studies have embedded IER detection methods in relatively long—and therefore potentially monotonous—questionnaires. In contrast, our data were collected in conjunction with an applied study, where great care was taken to provide a motivating experience to respondents in order to obtain valid responses. Thus, our findings may be limited in their generalisability to studies where using long test batteries is unavoidable.
